Papers with disease diagnosis
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines (2026.eacl-industry)
Copied to clipboard
Jean Seo, Gibaeg Kim, Kihun Shin, Seungseop Lim, Hyunkyung Lee, Wooseok Han, Jongwon Lee, Eunho Yang
| Challenge: | EPAG is a benchmark dataset and evaluation pipeline for pre-consultation of large language models. |
| Approach: | They propose a benchmark dataset and framework for evaluating pre-consultation ability of LLMs using diagnostic guidelines. |
| Outcome: | The proposed framework outperforms frontier LLMs in pre-consultation. |
A Layered Debating Multi-Agent System for Similar Disease Diagnosis (2025.naacl-short)
Copied to clipboard
| Challenge: | Traditional classification, contrastive learning, and large language models fail to detect subtle clues necessary for differentiation. |
| Approach: | They propose a framework that leverages Large Language Models to achieve accurate disease diagnosis . they structure patient information and integrate extensive medical knowledge to guide the analysis . |
| Outcome: | The proposed framework aims to identify subtle differences between similar diseases . the proposed framework can be used in clinical practice to improve accuracy . |
Benchmarking and Mitigating the Impact of Noisy User Prompts in Medical VLMs via Cross-Modal Reflection (2026.eacl-industry)
Copied to clipboard
| Challenge: | Existing medical vision-language models follow user-provided prompts blindly, a new study finds . current models are noisy, causing problems with reliability in real-world interactions . |
| Approach: | They propose a method to evaluate the influence of clinical prompts on medical vision-language models . they use cross-modal reflection chain-of-thought to train the model to produce reasoning paths . |
| Outcome: | The proposed method significantly improves the robustness against noisy prompts . existing Med-VLMs follow user-provided prompts blindly, the authors show . |
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models (2025.findings-naacl)
Copied to clipboard
Zhenyue Qin, Yu Yin, Dylan Campbell, Xuansheng Wu, Ke Zou, Ninghao Liu, Yih Chung Tham, Xiuzhen Zhang, Qingyu Chen
| Challenge: | Existing benchmarks for large vision-language models (LVLMs) are limited to ophthalmology-specific applications. |
| Approach: | They introduce a large-scale multimodal ophthalmology benchmark consisting of 21,993 instances across five ocular imaging modalities and 13 state-of-the-art LVLM representatives from closed-source, open-source and medical domains. |
| Outcome: | The proposed model shows significant performance drop in ophthalmology compared to other domains. |
CoAD: Automatic Diagnosis through Symptom and Disease Collaborative Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Automated diagnosis (AD) is a critical application of AI in healthcare . despite its simplicity and superior performance, a decline in disease diagnosis accuracy is observed . |
| Approach: | They propose a new collaborative disease and symptom generation framework to improve automatic diagnosis. |
| Outcome: | The Transformer-based method achieves an average 2.3% improvement over previous state-of-the-art methods . it can be used to query patients about their symptoms and health concerns . |
Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing models struggle to balance predictive accuracy with human-understandable rationales. |
| Approach: | They propose to enhance LLMs by leveraging rationale distillation and domain knowledge injection for trustworthy multimodal rationale generation. |
| Outcome: | Experiments on real-world medical datasets show that ClinRaGen achieves state-of-the-art performance in disease diagnosis and rationale generation. |
RareSyn: Health Record Synthesis for Rare Disease Diagnosis (2025.emnlp-main)
Copied to clipboard
| Challenge: | RareSyn is a data synthesis approach to augment and de-identify EHRs with a focus on rare diseases. |
| Approach: | They propose a data synthesis approach to augment and de-identify EHRs with a focus on rare diseases. |
| Outcome: | The proposed model augments and de-identifies EHRs with a focus on rare diseases. |
MKeCL: Medical Knowledge-Enhanced Contrastive Learning for Few-shot Disease Diagnosis (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing approaches to disease classification are limited in real-world clinics due to insufficient data and inflexibility. |
| Approach: | They propose a medical knowledge-Enhanced Contrastive Learning approach to disease diagnosis . they incorporate medical knowledge graphs and medical licensing exams in modeling . |
| Outcome: | The proposed model outperforms existing models on real clinical EMRs on a single patient. |
DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaboration (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing LLMs fail to capture the dual nature of medical consultation (MC) this mismatch often results in ineffective symptom inquiry and unreliable disease diagnosis. |
| Approach: | They propose a novel LLM-based framework that performs Dual-Decision Optimization by decoupling the two sub-tasks and optimizing them with distinct objectives through a collaborative multi-agent workflow. |
| Outcome: | The proposed framework outperforms existing LLM-based approaches on three real-world MC datasets and achieves competitive performance with state-of-the-art generation-based methods. |
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent Driven Chain-of-Inquiry (2026.findings-acl)
Copied to clipboard
Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin, Hasan Muhammad Abdullah, Md Mehedi Hasan, Mohammad Zabed Hossain, Shifat E. Arman
| Challenge: | Visual question-based reasoning is a key component of vision-language models. |
| Approach: | They propose a framework for visual question-answering that integrates visual intent with visual severity to improve diagnostic accuracy. |
| Outcome: | The proposed framework improves diagnostic correctness, reduces hallucination, and increases reasoning efficiency. |